Papers with multi-armed bandit learning
Multi-Source Test-Time Adaptation as Dueling Bandits for Extractive Question Answering (2023.acl-long)
Copied to clipboard
| Challenge: | Recent research on test-time adaptation suggests a possible way to improve the generalization ability of LLMs. |
| Approach: | They propose to use multi-armed bandit learning and multi-arm dueling bandits to solve a multi-source test-time model adaptation problem from user feedback. |
| Outcome: | The proposed model is more effective than other strong baselines on extractive question answering datasets. |